Papers with speech resynthesis

2 papers
textless-lib: a Library for Textless Spoken Language Processing (2022.naacl-demo)

Copied to clipboard

Challenge: Textless spoken language processing is an exciting area of research that promises to extend applicability of the standard NLP toolset onto spoken language and languages with few or no textual resources.
Approach: They introduce textless-lib, a PyTorch-based library that provides textless spoken language processing tools.
Outcome: The proposed library significantly simplifies research in the textless setting and will be a handful for speech researchers and the NLP community at large.
EmphAssess : a Prosodic Benchmark on Assessing Emphasis Transfer in Speech-to-Speech Models (2024.emnlp-main)

Copied to clipboard

Challenge: EmphAssess evaluates speech-to-speech models' ability to encode and reproduce prosodic emphasis across a change of speaker and language.
Approach: They propose a prosodic benchmark to evaluate the ability of speech-to-speech models to encode and reproduce prosodic emphasis.
Outcome: The proposed model can encode and reproduce prosodic emphasis across speech inputs and outputs . EmphaClass classifies emphasis at the frame or word level .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations